Tags: topic: financial technology*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Zhening Li and colleagues introduce JAZ, an LLM agent framework designed to minimize the complexity of agent loops by treating them as a programming language primitive called `invoke`. Instead of relying on external specialized systems for memory or self-improvement, JAZ enables agents to achieve these capabilities through code execution where all interactions are treated as variables within the environment. This minimalist approach allows highly expressive workflows, such as long-horizon recall and continual self-improvement, using only prompting rather than manually designed tools or complex architectures.

    - The `invoke` primitive allows for recursive calls, enabling LLMs to write arbitrary executable code that includes further iterations of itself.
    - In testing on the StuLife dataset, JAZ outperformed MemGPT (Letta) by 8% in recall performance while costing half as much.
    - On self-improvement tasks using AppWorld, JAZ demonstrated a 4% improvement over ACE at a lower computational cost.
  2. Convai Innovations presents Laya, a multilingual, non-autoregressive system 1 decision model designed to provide typed answers with mathematically calibrated probabilities in a single forward pass. Unlike generative models, it does not generate text, thereby eliminating hallucinations and the need for parsing. The framework includes an automated Router that detects language and script to dispatch tasks to the most efficient checkpoint (English or Multilingual) within approximately 35ms on GPU.
    - It is trained using Reinforcement Learning with Calibrated Decisions (RLCD) to ensure honest probability reporting.
    - Laya can support context lengths of up to 8,192 tokens in its multilingual version.
    - The model family includes specialized checkpoints like `laya-typed-decisions` which achieves significantly higher accuracy through fine-tuning on specific workflows.
    - Performance benchmarks show it is roughly 6–8× faster than TypeSafe Jev for single question latency on a T4 GPU.
  3. Diogo Almeida writes that TypeSafe AI is releasing Jev, its first System One Model—a new class of frontier model built for fast, structured decisions that software can consume directly. Unlike autoregressive language models that generate strings token by token, Jev outputs type-safe structured values with calibrated probabilities in a single parallel query, achieving frontier-level intelligence on decision tasks at roughly 40–200× lower latency and cost. The company's new training method, Reinforcement Learning for Calibrated Decisions (RLCD), optimizes for epistemically honest probability estimates rather than human preference or verifiable rewards, and the architecture is mathematically incapable of producing type errors or hallucinations.
    - Named after William Stanley Jevons, whose paradox predicted that efficiency gains would increase (not decrease) total demand; TypeSafe expects each order-of-magnitude cost drop to unlock orders of magnitude more use cases.
    - Workflow evals benchmark Jev against the average of GPT-6 Astra and Fable 5.1 as reference probabilities, claiming 193.6× speed and 444.6× cost advantages on production-shaped tasks.
    - The team demonstrated real-time intelligence with a Doom bot making 10 structured queries per second (~$7/hour) and a Wikiracing bot that outperforms LLMs at high-cardinality link selection.
    - Jev supports output cardinality up to 255; for higher-cardinality choices it falls back to a two-stage scoring system that scores independently then makes an explicit selection.
  4. Milan Minsky writes that Leela AI transforms standard factory and warehouse cameras into smart sensors, offering an alternative to traditional IoT sensors by leveraging existing video feeds instead of physical hardware. The platform provides contextual visibility into operations, identifies bottlenecks, and tracks interactions between machines, operators, and materials without requiring retrofitting. It complements IoT systems by integrating with platforms like Velotic ThingWorx and AVEVA to create a comprehensive digital twin of manufacturing floors. The core technology utilizes MIT research-based AI, combining causal and neural networks for efficient data processing.
  5. Dan Russell writes about the power of AI-augmented search to retrieve hard-to-find information, using an example of finding a study on how the gender of lab assistants affects experimental outcomes on lab mice. He demonstrates how a simple query with AI can yield relevant results, leading to original source papers. The study highlights the impact of experimenter gender on reproducibility in scientific research.
  6. El Assadi et al. compare ten LLMs (six families) and 26 embedding models (118M - 14B parameters) on 37 tasks, considering cost. In aggregate, the two paradigms are effectively tied (best LLM scores 77.6 versus best embedding model 77.2), yet their strengths diverge by task: LLMs lead on reasoning-heavy retrieval while embedding models lead on classification, and the two match on clustering, STS, and pair classification.

    LLMs are significantly more expensive (up to 1,431x) and slower (2.5-736x) than embedding models for certain tasks. The authors suggest using embedding models for similarity, classification, and clustering, and LLMs for reasoning in retrieval.
    Reasoning tokens are 28-81% of LLM inference cost; lower budgets maintain or boost retrieval quality for most tested models.
    - Only Gemini 3.1 Pro breaks into the Pareto frontier alongside the leading embedding models.
    - Accepted to COLM 2026; code, datasets, and results are publicly released on GitHub.
  7. Ashish Vaswani et. al. introduce Transformers and Attention in this classic 2017 paper.

    The Transformer architecture relies solely on attention mechanisms, dispensing with recurrence and convolutions entirely for sequence transduction tasks. This new network design improves translation quality while being more parallelizable and significantly faster to train than previous models.

    - Achieved 28.4 BLEU on the WMT 2014 English-to-German translation task.
    - Reached a state-of-the-art score of 41.8 BLEU for English-to-French using eight GPUs in only 3.5 days.
    - Demonstrates successful application to English constituency parsing with both large and limited training data sets.
  8. @0xabad1dea@infosec.exchange writes about an incident where AI-assisted mathematical proofs appear to exploit bugs in theorem provers, specifically highlighting a case involving the Collatz conjecture and Lean 4. The discussion explores whether large language models are inadvertently discovering software vulnerabilities through pattern matching or learning from existing technical discussions about those bugs, while broader debates address the inherent limitations of formal verification when facing hardware faults, modeling errors, and human mistakes in specifications.
  9. The community-led open-source hosting site Codeberg has announced bans on two types of projects: cryptocurrency-related projects and those whose code is substantially or entirely generated by Large Language Models (LLMs) such as Claude or OpenAI Codex. Following a community vote, the ban on LLM-generated code passed with 358 votes in favor to 144 against. The reasoning for these decisions includes concerns over "license whitewashing," the massive increase in hardware and energy costs caused by AI datacenter scaling, and the potential negative impact of generative AI tools on the Open Source Software (OSS) community.

    The comments reflect a deep division within the tech community regarding this decision:
    * **Supporters** argue that current LLM practices are unethical because they undermine software rights, increase environmental strain, and create massive amounts of "junk" code that is difficult to maintain or scale.
    * **Critics/Skeptics** suggest the ban is a "Luddite" reaction to an unstoppable trend (comparing it to people refusing cell phones). They argue that LLMs are already integrated into most workflows ("the toothpaste is out of the tube") and that banning them might be impossible or impractical.
    * **Nuanced Perspectives** emerge from users who distinguish between using LLMs as a "reasoning tool" for scientific/mathematical scaffolding versus pure "vibe coding." Some argue that while full generation creates maintenance risks, LLM tools are essential assets for hobbyists and professionals alike to solve problems efficiently.
  10. An exploration into the history of conversational technology, tracing its roots from Joseph Weizenbaum's 1966 ELIZA experiment at MIT to modern large language models like ChatGPT and Claude. The article examines how the evolution from rule-based symbolic AI to probabilistic deep learning has changed human interaction with machines, often leading users to attribute human qualities to code. It specifically addresses the risks of "chatbot psychosis" and the danger of individuals relying on general-purpose generative models for mental health support when these systems are prone to hallucinations or reinforcing delusional beliefs.

    * The transition from symbolic AI's explicit rules to modern deep learning
    * Joseph Weizenbaum’s warning against humanizing machines via the ELIZA effect
    * The psychological impact and risks of using large language models for emotional support

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "topic: financial technology"

About - Propulsed by SemanticScuttle